Skip to content

ateom: remove the ateom directory on graceful shutdown - #1678

Merged
Jeff Luo (JeffLuoo) merged 3 commits into
agent-substrate:mainfrom
baizhenyu:ateom-dir-self-cleanup
Sep 19, 2026
Merged

Jeff Luo (JeffLuoo) merged 3 commits into
agent-substrate:mainfrom
baizhenyu:ateom-dir-self-cleanup

Conversation

@baizhenyu

@baizhenyu Tim Bai (baizhenyu) commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

Part of #1677 — the graceful-shutdown half; the atelet-side janitor for ungraceful exits is tracked there.

Nothing removes ateoms/<podUID>/ when a worker pod goes away, so every pool rollout leaks one directory (holding a dead ateom.sock) per replaced worker — and since the stats sweep enumerates that directory as its discovery registry (#961), each leak costs a dial+probe per sweep, forever.

Change

A defer os.RemoveAll(ateomDir) in both runtime mains, registered right after the MkdirAll that creates the directory. ~7 lines each plus the comment carrying the safety argument.

Why this is safe:

  • Same-pod container restart: kubelet serializes containers — the replacement is not started until this process, and therefore this defer, has exited. The remove cannot race the successor's MkdirAll.
  • Shutdown ordering: the SIGTERM path drains, GracefulStops, Serve returns, do returns — the defer runs after the listener is already closed, so nothing can be dialing the socket it removes.
  • Boot-failure paths: error returns before Listen also run the defer; the pod restarts and re-creates the directory. Harmless.

What it deliberately does not cover: SIGKILL after the grace period, OOM kills, node crashes — defers don't run there by nature. That residue is the janitor's job (#1677); this change shrinks the janitor's caseload to exactly those, since rollouts — the dominant growth driver — terminate gracefully.

Testing

Both runtime mains are untested boot plumbing (no unit seam exists for do()); verified by linux builds and test-compiles of both packages.

Live-verified on the ate-dev GKE cluster (gVisor pool; the micro-VM change is the identical lines):

  • Baseline census of one node found 11 ateom directories with exactly 1 live pod — ten stale leftovers from prior rollouts, the leak in the wild.
  • Rolled the pool onto an image built from this branch. The outgoing pre-change pod (445af5aa…) leaked its directory on termination, reproducing the bug side-by-side with the fix.
  • Gracefully deleted a pod running this branch (5618282c…): its directory was removed on shutdown, while the pre-change pod's leaked directory remained in the same census — before/after behavior on the same node, same listing.

Nothing removes ateoms/<podUID>/ when a worker pod goes away, so every
pool rollout leaks one directory per replaced worker onto the node --
and since the stats sweep enumerates that directory as its discovery
registry, each leak costs a dial+probe per sweep forever (agent-substrate#1677).

A defer after the MkdirAll covers the exits that dominate that growth:
rollouts terminate gracefully, so the SIGTERM path (drain, GracefulStop,
Serve returns, do returns) runs it, as do error returns during boot.
Safe against a same-pod container restart because kubelet does not
start the replacement container until this process -- and so this
defer -- has exited. Ungraceful exits (SIGKILL, OOM, node crash) skip
defers by nature; the atelet-side janitor proposed in agent-substrate#1677 is the
backstop for those, and this change shrinks its caseload to them.

Part of agent-substrate#1677.
Comment thread cmd/ateom-gvisor/main.go
An empty or path-shaped pod UID was previously harmless -- MkdirAll
just created a wrong path -- but the shutdown RemoveAll turned it
destructive: AteomPath("") is AteomsDir() itself, so a misconfigured
ateom would delete every ateom's socket directory on the node. Fail
fast at boot instead, with the validator the atelet RPC boundary
already uses.
Comment thread cmd/ateom-gvisor/main.go
os.Exit(1) on a Serve failure skipped every defer in do(), including
the tracer and meter shutdowns and now the ateom directory removal.
Return the error the way ateom-microvm does; main logs it and exits.
@JeffLuoo
Jeff Luo (JeffLuoo) added this pull request to the merge queue Sep 18, 2026
Merged via the queue into agent-substrate:main with commit 27bf344 Sep 19, 2026
14 of 15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/gvisor area/microVM area/node kind/bug Something isn't working / bugfixes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants